Google EmbeddingGemma 2 adds multimodal sentence transformers

6 10 2026
Google EmbeddingGemma 2 adds multimodal sentence transformers

Google has released EmbeddingGemma 2, a sub-1 billion parameter model that maps text, code, images, video, and audio into a shared 768-dimensional space. Integrated into sentence-transformers v6.1.0, the architecture is modular so you can drop unused modality encoders at load time to scale the footprint from 270M parameters up to 740M. It includes support for task-specific prompts and native dimension truncation down to 128 dimensions.

If you are building a local RAG pipeline or cross-modal search tool that needs to handle mixed media without a heavy inference server, this is worth testing. The catch is that matching modalities still requires strict adherence to prompt names and specific token markers, and sub-1B models will inevitably lag behind larger proprietary endpoints on nuanced retrieval tasks. Skip it if your stack is purely text-based and already tuned around existing dense retrievers.

more: https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/


Actions

Information

Leave a comment